Why test data management is essential for SOC 2 compliance 

A stock image showing someone working on a laptop
Comments 0

Share to social media

Using live data outside production is one of the fastest ways to create compliance risk. It becomes harder to control who can access the data, how it’s handled, and how long it’s kept for. However, it doesn’t have to be that way. Enter Test Data Management (TDM).

A TDM approach provides the controls SOC 2 auditors look for in this situation. It’s an automated, traceable, end-to-end process for protecting, provisioning, and removing customer data so that it can be used safely in non-production environments. 

First, let’s answer the big question: what exactly is SOC 2 compliance, and why is it so important?

What is SOC 2 compliance, and why is it so important? 

SOC 2 is designed to help answer a simple question that any customer must ask of a service provider: can we trust you with our data?  

SOC 2 compliance isn’t a legal requirement in the way GDPR or HIPAA can be. However, as more organizations move applications into the cloud and rely on SaaS (software-as-a-service) providers, it’s become a standard way for them to check that their chosen providers handle sensitive customer information safely.

Many larger enterprises – especially in the U.S. market – won’t engage SaaS, cloud, or B2B technology vendors unless they can provide a SOC 2 Type II report. This report assesses not only that appropriate controls exist, but whether they operate effectively over time. 

“SOC 2 evaluates whether an organization’s systems and controls effectively protect and manage customer information” — AICPA Trust Services Criteria 

What does SOC 2 require?

The System and Organization Controls (SOC) framework was developed by the American Institute of Certified Public Accountants (AICPA), and SOC 2 is an independent attestation report based on their Trust Services Criteria (TSC). It establishes the sort of controls an organization needs to demonstrate for security, availability, processing integrity, confidentiality, and privacy. 

The Security category is mandatory for every SOC 2 report. It assesses an organization’s system and data security controls against a set of Common Criteria (CC1–CC9), covering: 

  • CC1–CC5: The foundations of control 
    How an organization’s controls are owned, managed and monitored. 
  • CC6: Access control and data handling 
    Controlling access to systems and data; identifying and protecting sensitive data. 
  • CC7: System operations 
    Monitoring for, detecting, and responding to security incidents. 
  • CC8: Change Management 
    Ensuring all system updates are authorized, documented and properly tested. 
  • CC9: Risk mitigation 
    Identifying threats and planning for operational resilience. 

The other four categories are only included if they match your service commitments. If you handle proprietary business data or personally-identifiable information (PII), for example, then Confidentiality (C) or Privacy (P) applies to ensure secure handling and disposal.

Similarly, Availability (A) may be included for specific ‘uptime’ or 24/7 availability promises, while Processing Integrity (PI) applies where you make commitments about the accuracy, completeness, and timeliness of data processing outputs (for example, for financial transactions or data analytics). 

Across all categories, SOC auditors don’t just want to see the right technologies, such as encryption or access controls. Instead, they expect documented, repeatable processes for data handling and protection – plus evidence that these processes are followed consistently and kept under review.

Evidence might include logs, reports, and change records showing controls are operating as intended and are being updated as systems, and requirements, change. 

What are the SOC 2 challenges that test data management (TDM) helps to solve?

The rest of this article focuses on where test data management (TDM) most directly supports SOC 2: 

  • Protecting data outside production (CC6 plus Confidentiality and Privacy criteria) 
    Production systems are often tightly controlled, but SOC 2 also expects customer data to be protected before it’s copied, accessed, and reused outside production for development, testing, analytics, or AI workflows. 
  • Supporting safe, realistic testing for change management (CC8) 
    Test data must always remain useful for its intended purpose without exposing sensitive information. This is so it can be used safely to test changes and catch regressions before release.

Teams that continue to rely on manual processes to identify, protect, and provision test data will struggle with both of these challenges. In fact, according to Redgate’s 2026 State of the Database Landscape report, as many as 39% of the survey’s respondents fall into that category.

“Despite growing investment in modern delivery practices, many organizations are still relying on manual processes to test and deploy database changes”

– The State of the Database Landscape 2026 report
Learn more

These manual processes often lead to incomplete or unrealistic datasets, making it harder to test changes with confidence or trust analytical results. 

By contrast, a TDM approach supports SOC 2 expectations by embedding data protection into the test data lifecycle while keeping data realistic and fit for purpose. It standardizes and automates the identification of sensitive and personal data, ensuring data is always protected before it moves outside production.

It then automates data provisioning and cleanup as a controlled workflow, with traceable records of what was done, where data went, and when it was removed. 

How does TDM protect customer data outside production? 

“The entity identifies and maintains the confidentiality of information designated as confidential from its receipt or creation through its retention and disposal.” (C1.1, C1.2) 

Woven throughout SOC 2’s Trust Services Criteria is the expectation that confidential and personal data will be protected across its lifecycle. This means from creation all the way through to retention and use. It also includes reuse outside of production, and secure removal when it’s no longer needed for its intended purpose.

In practice, it means an organization must first define and identify confidential/personal information (C1.1), and then apply controls that: 

  • Limit use of personal information to the identified purposes (P4.1) 
  • Retain it only as long as necessary and dispose of it securely (P4.2, P4.3) 
  • Maintain a record of detected or reported unauthorized disclosures (P6.3) 
  • Classify sensitive information by its relevant characteristics (CC2.1 points of focus) 
  • Use logical access controls to restrict use of protected information, based on an inventory of information assets (CC6.1) 
  • Restrict transmission, movement, and removal of sensitive data, and protect it during handling (CC6.7) 
  • Test system changes while ensuring sensitive data remains protected (CC8.1) 

This is a challenging part of SOC 2 compliance for any enterprise. Once customer data leaves production, different teams access it for different reasons, and safe handling controls become harder to enforce without a consistent, automated way to apply them across every environment.

A TDM approach supports SOC2 expectations by embedding controls into a repeatable workflow for identifying sensitive data, protecting it before reuse, and controlling how sanitized copies are provisioned and cleaned up. 

How a TDM approach identifies and classifies personal and sensitive data (C1.1, CC2.1, CC6.1) 

In most database estates, sensitive data rarely comes neatly labelled. It’s spread across many tables, is often duplicated into reporting structures, and can hide behind obscure column names. It can also appear unexpectedly in free-text fields and attachments. 

Manual classification is slow, labor-intensive, and error-prone. It also becomes unrealistic at scale, as data volumes grow and schemas evolve. 

A TDM approach, by contrast, fully supports SOC 2’s expectation to classify information by relevant characteristics (CC2.1), and to apply access controls based on an inventory of information assets (CC6.1). It will: 

  • Automate discovery and classification, often incorporating AI-assisted identification.
  • Provide a predefined taxonomy – but allow teams to extend and customize it without adding brittleness or complexity. 
  • Give teams a single source of truth for data protection before any copies are distributed to non-production environments.

How a TDM approach ensures consistent data protection before reuse (CC6, CC8) 

Once sensitive data has been identified, SOC 2 expectations shift from visibility to control. Organizations need to restrict access to sensitive data (CC6.1), restrict its movement wherever possible (CC6.7), and protect it during development, testing, and change processes (CC8.1). 

The ‘old ways’ of providing test data generally do not meet these compliance challenges. Copying live data to test environments creates PII exposure (even on a secure shared server), and relies on proper data handling practices from each team.

Then there’s manual sanitization techniques that require writing and maintaining often-complex masking scripts, which are unreliable and don’t scale well. These scripts can miss PII, produce different results across environments, and leave little evidence of what was protected, and when.

They are also brittle and hard to keep in step with schema changes. 

A TDM approach replaces this with a controlled, repeatable workflow that produces a sanitized base copy of the data that can then be distributed for approved non-production uses. The data protection process will use one or more of the following techniques: 

  • Static data masking – irreversible replacement of sensitive values while preserving referential integrity and realistic data distributions. 
  • Data minimization using subsetting – minimal but representative datasets, excluding data that is not required for the intended use. This reduces both exposure and operational overhead. 
  • Synthetic data generation – using statistical/rule-based or AI-based techniques to create datasets that mimic production’s characteristics and distributions without containing any real records. This is especially valuable for highly sensitive data domains, or for generating edge-case and failure conditions that don’t occur naturally in a production snapshot.

Great, but what makes this process maintainable and scalable?

Simplicity and automation are what make this process maintainable and scalable. If a masking run needs a custom dataset for a specialized data type, or a column-specific generator to satisfy complex constraints, these should be added as configuration options rather than manually-applied custom scripts that teams then must maintain. 

When classifications and protection definitions are stored as simple configuration, tracked in version control, the process becomes easier to automate, maintain and adapt as systems change. There’s also far less reliant on ad-hoc scripts. 

Critically, it also ensures rules are applied based on the same classifications each time data is provisioned. This keeps the protection consistent across environments. 

Move fast. Govern at scale.

Redgate Flyway Enterprise embeds guardrails in the database layer, so every change is policy-checked, deterministic, and traceable.
Try for free

How a TDM approach ensures controlled, auditable test data provisioning and cleanup (P4.1, P4.2, P4.3) 

SOC 2’s Confidentiality and Privacy criteria reinforce the simple expectation that customer data must stay protected wherever it’s used.

Of course, this isn’t just about anonymization; no compliant data protection strategy can rely on one technique in isolation. It also covers purpose limitation, access restriction, retention, and disposal. SOC auditors then assess how well these controls hold up across every environment where customer data is used.

This is another area where manual processes struggle, since it’s difficult to produce a clear record of what data was provisioned, where it went, who had access to it, and when it was removed. Auditors generally look for systemic controls that operate consistently over time, rather than relying on manual steps. 

TDM practices tackle this by treating provisioning and cleanup as part of the same automated and controlled workflow: 

  • Automated provisioning and cleanup rules. Define where sanitized datasets can be provisioned, how long they can be retained, and when they must be removed. This often happens by provisioning test databases as short-lived containerized or virtualized clones – created, refreshed, and torn down on a controlled schedule. This reduces sprawl, limits the ‘attack surface’, and supports retention and disposal expectations. 
  • Strictly controlled access to data provisioning. Restrict who can access provisioning/refresh workflows using centralized authentication (for example, OpenID Connect, or OIDC). 
  • Traceable definitions and repeatable execution. Store classifications and masking definitions as configuration. Keep an audit trail of what ran, when it ran, and what rules were applied. 
  • End-to-end auditability. Record how data was classified, protected, provisioned, accessed, and removed. This supports the “ongoing operational effectiveness” nature of SOC 2 type II reporting. 

How does a TDM approach support SOC 2 testing expectations?

“The entity authorizes, designs, develops or acquires, configures, documents, tests, approves, and implements changes to infrastructure, data, software, and procedures to meet its objectives.” – CC8.1 

Most database environments are changing at an increasingly rapid pace, driven by schema updates, application releases, patching, and configuration changes. Some changes are planned upgrades; others are urgent fixes in response to incidents. SOC 2 expects all of these changes to be tested properly, with sensitive data staying protected throughout.

Of course, this only works if teams can get test data that is safe to reuse and still behaves like the real thing – exactly where many test data strategies fail. If data protection produces datasets that are inconsistent, incomplete, or unrealistic, testing becomes unreliable.

The result of unreliable testing? Lower test coverage, more manual checks, and greater risk during releases and incident fixes. 

A TDM approach will provide realistic datasets without exposing customer data. Relationships are maintained, data distributions stay realistic, and the same test dataset can be recreated consistently as changes are tested and retested. 

How a TDM approach ensures system changes are tested thoroughly but safely (CC8.1) 

A TDM approach can provide test datasets required to support all the types of tests referenced by SOC 2, in CC8.1.  

Automated and controlled preparation of test databases 

Redgate Flyway Enterprise can help here. It can automatically create a target test database at the current schema version, generate the SQL changes under test, and check them deterministically for code policy violations.

It can then automatically load the designated test dataset and deploy the new version securely to the test environment ready for the test case. 

Type of test Example Test data requirements 
Unit testing Validate a small change to a stored procedure or function.Small, purpose-built datasets that typically only need data that the object can reference directly. 
Integration and regression testing Verify an application workflow still works after a schema/API (application programming interface) change.Realistic, immutable datasets where the result can be cross-checked by the business for validity. 
User acceptance / QA testing Validate behavior end-to-end before release.Production-like datasets that reflect real workflows and edge cases without PII exposure.
Patch and update testing Test database engine or application patches before rollout. Representative data volumes and distributions so that performance and query plans reflect production. 
Bug fix/validation Reproduce a production issue, test the fix, and confirm no new issues. A dataset that mirrors the failure conditions, delivered safely and repeatably.

How a TDM approach supports incident recovery and safely ensures resilient testing (CC7.5, CC8.1) 

CC7.5 expects organizations to have a documented incident recovery plan and to test it on a regular basis. These recovery tests are only meaningful when environments reflect real data volumes, relationships, and workload patterns. 

TDM supports this by providing production-like datasets that are safe to reuse. With this, teams can rehearse recovery procedures and validate outcomes without distributing raw customer data into non-production.

It also supports the broader CC8.1 expectation that changes and recovery procedures can be tested safely as part of development and change processes. 

In conclusion: why you should utilize test data management for SOC 2 compliance

SOC 2 is not just about whether an organization has documented controls for protecting customer data. Auditors look for evidence that those controls operate consistently over time, meaning they are actively maintained and built into everyday database work rather than relying on manual intervention. 

They’ll also expect to see those controls extending to any copies or derivatives of that data reused outside production, for development and testing, incident fixes, and analytics. This is where a TDM approach provides a consistent, automated way to meet these expectations. It will: 

  • Create safe, representative test datasets through data masking, subsetting, and data generation. 
  • Provision and refresh test data through an automated workflow, coordinated through version-controlled configuration. This ensures consistent enforcement of data access, movement, retention, and cleanup across environments. 
  • Maintain traceable evidence that proves how sensitive information is protected throughout its lifecycle. 

A test data management (TDM) approach allows you to test changes more thoroughly, using realistic data without risking exposure of customer information. This improves the quality and repeatability of testing while supporting SOC 2’s expectation that sensitive data stays protected throughout its lifecycle.

Simple Talk is brought to you by Redgate Software

Take control of your databases with the trusted Database DevOps solutions provider. Automate with confidence, scale securely, and unlock growth through AI.
Discover how Redgate can help you

FAQs: Test data management for SOC 2 compliance

1. Is test data management required for SOC 2 compliance?

No. SOC 2 doesn’t name TDM directly, but it does require controls for protecting data outside production (CC6), safe testing during change management (CC8), and secure retention and disposal (Privacy criteria). TDM is one of the most direct ways to meet those requirements.

2. What SOC 2 criteria does test data management support?

Mainly CC6 (access control and data handling), CC8 (change management), and the Confidentiality and Privacy criteria (C1.1, P4.1–P4.3). It also supports CC7.5 for incident recovery testing.

3. Can I use production data for testing under SOC 2?

Not safely. Copying live data into test environments creates PII exposure and depends entirely on every team handling it correctly. SOC 2 auditors expect sanitized, access-controlled alternatives instead.

4. What's the difference between data masking and synthetic data generation?

Masking irreversibly replaces sensitive values in real data while keeping it realistic. Synthetic data generation creates entirely new datasets that mimic production patterns without containing any real records. Both count as valid TDM techniques.

5. Does test data need to be deleted after use for SOC 2?

Yes. SOC 2’s Privacy criteria (P4.2, P4.3) expect data to be retained only as long as necessary and disposed of securely. Automated provisioning and cleanup workflows are the standard way to demonstrate this.

This document contains proprietary information and is protected by copyright law.

Copyright © 2026 Red Gate Software Limited. All rights reserved

Article tags

About the author

Tony Davis

See Profile

Tony Davis is an Editor with Red Gate Software, based in Cambridge (UK), specializing in databases, and especially SQL Server. He edits articles and writes editorials for both the Simple-talk.com and SQLServerCentral.com websites and newsletters, with a combined audience of over 1.5 million subscribers. You can sample his short-form writing at either his Simple-Talk.com blog or his SQLServerCentral.com author page.

As the editor behind most of the SQL Server books published by Red Gate, he spends much of his time helping others express what they know about SQL Server. He is also the lead author of the book, SQL Server Transaction Log Management.

In his spare time, he enjoys running, football, contemporary fiction and real ale.